PLOS Digital Health
● Public Library of Science (PLoS)
All preprints, ranked by how well they match PLOS Digital Health's content profile, based on 106 papers previously published here. The average preprint has a 0.26% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Huang, Y.; Sharma, P.; Palepu, A.; Greenbaum, N.; Beam, A. M.; Beam, K.
Show abstract
ImportanceArtificial intelligence (AI) based on deep learning has shown promise in adult and pediatric populations in the interpretation of medical imaging to make important diagnostic and management recommendations. However, there has been little work developing new AI methods for neonatal populations. ObjectiveTo develop a novel, deep contrastive learning model to predict a comprehensive set of pathologies from radiographs relevant to neonatal intensive care. Design, Setting, and ParticipantsWe identified a retrospective cohort of infants who obtained a radiograph while admitted to a large neonatal intensive care unit in Boston, MA from January 2008 to December 2023. After collecting radiographs with corresponding reports and relevant demographics for all subjects, we randomized the cohort into three sets: training (80%), validation (10%), and test (10%). InterventionsWe developed a deep learning model, NeoCLIP, to identify 15 unique pathologies and 5 medical devices relevant to neonatal intensive care on plain film radiographs. The pathologies were automatically extracted from radiology reports using a custom pipeline based on large language models. Main Outcomes and MeasuresWe compared the performance of our model, as defined by AUROC, against various baseline methods. ResultsWe identified 4,629 infants which were randomized into the training (3,731 infants), validation (419 infants), and test (479 infants) sets. In total, we collected 20,154 radiographs with a corresponding 15,795 reports. The AUROC of our model was greater than all baseline methods for every radiographic finding other than portal venous gas. The addition of demographics improved the AUROC of our model for all findings, but the difference was not statistically significant. Conclusions and RelevanceNeoCLIP successfully identified a broad set of pathologies and medical devices on neonatal radiographs, outperforming similar models developed for adult populations. This represents the first such application of advanced AI methodologies to interpret neonatal radiographs.
Sangeda, R. Z.; Masamu, U.; Kandonga, D.; Mbuya, F.; Mwita, L.; Makundi, F.; Mgaya, J.; Osati, E. F.; Jonathan, A.; Mmbando, B. P.; Nkya, S.; Balandya, E.; Makani, J.
Show abstract
ObjectiveTo document the deployment and training of Research Electronic Data Capture (REDCap) at Muhimbili University of Health and Allied Sciences (MUHAS) in Tanzania and to evaluate user satisfaction and perceived impact on research practices in Tanzania. Materials and MethodsSix structured REDCap training workshops were conducted between 2016 and 2022. The post-training evaluation surveys were completed by 110 participants. Descriptive statistics were generated for demographics and cadres, while Likert-scale responses assessed satisfaction and perceived utility. Diverging stacked bar charts were used to visualize attitudes toward adopting REDCap. ResultsMost trainees were male (62.7%), aged 26-32 years (52.7%), and affiliated with MUHAS (66.4%), with additional representation from Ghana and Nigeria. The participants included academic or research staff (41.8%), postgraduate students (28.2%), undergraduate students (11.8%), and other professionals (18.2%). Post-training evaluations indicated consistently high satisfaction, with mean scores above 4.4 on a 5-point scale. Trainees strongly endorsed REDCaps ability to enhance research effectiveness (mean 4.53, SD 0.65), increase productivity (mean 4.52, SD 0.69), and improve task completion (mean 4.50, SD 0.68). More than 85% of the respondents agreed or strongly agreed with the positive statements, underscoring the broad acceptance of REDCap as a reliable tool for research data management. DiscussionThe findings underscore the importance of structured training in mitigating barriers to adoption and enhancing data quality, efficiency, and institutional visibility. ConclusionStructured REDCap training in Tanzania was associated with high user satisfaction and perceived improvements in research productivity and data quality, offering a scalable model for strengthening data management capacity in sub-Saharan Africa.
Lee, J.; Xu, X.; Kim, D.; Deng, H. H.; Kuang, T.; Lampen, N.; Fang, X.; Gateno, J.; Yan, P.
Show abstract
PurposeThis study examines the application of Large Language Models (LLMs) in diagnosing jaw deformities, aiming to overcome the limitations of various diagnostic methods by harnessing the advanced capabilities of LLMs for enhanced data interpretation. The goal is to provide tools that simplify complex data analysis and make diagnostic processes more accessible and intuitive for clinical practitioners. MethodsAn experiment involving patients with jaw deformities was conducted, where cephalometric measurements (SNB Angle, Facial Angle, Mandibular Unit Length) were converted into text for LLM analysis. Multiple LLMs, including LLAMA-2 variants, GPT models, and the Gemini-Pro model, were evaluated against various methods (Threshold-based, Machine Learning Models) using balanced accuracy and F1-score. ResultsOur research demonstrates that larger LLMs efficiently adapt to diagnostic tasks, showing rapid performance saturation with minimal training examples and reducing ambiguous classification, which highlights their robust in-context learning abilities. The conversion of complex cephalometric measurements into intuitive text formats not only broadens the accessibility of the information but also enhances the interpretability, providing clinicians with clear and actionable insights. ConclusionIntegrating LLMs into the diagnosis of jaw deformities marks a significant advancement in making diagnostic processes more accessible and reducing reliance on specialized training. These models serve as valuable auxiliary tools, offering clear, understandable outputs that facilitate easier decision-making for clinicians, particularly those with less experience or in settings with limited access to specialized expertise. Future refinements and adaptations to include more comprehensive and medically specific datasets are expected to enhance the precision and utility of LLMs, potentially transforming the landscape of medical diagnostics.
Dhaimade, P. A.; Henderson, R.
Show abstract
BackgroundLarge language models (LLMs) have demonstrated rapid advancements in natural language understanding and generation, prompting their integration into biomedical research, clinical practice, and professional education. However, systematic evaluation of LLMs in specialty-specific domains such as dentistry and periodontology remain limited, particularly regarding multidimensional performance metrics. ObjectiveTo conduct a comprehensive, multidimensional assessment of commercially available LLMs: GPT-4.0, GPT-5.0, and Claude SONNET 4.0 on the American Academy of Periodontology in-service examination, focusing on response accuracy, self-assessed confidence calibration, citation validity, and hallucination prevalence. MethodsModels were evaluated on the 2024 AAP In-Service Examination (331 questions) using two formats: Full Test (all questions at once) and Individual Question (one at a time). Prompts were standardized; models selected answers, and for GPT-5.0 and Claude SONNET 4.0, also provided confidence ratings and citations. Citation validity was assessed using a human-in-the-loop protocol with expert review. Statistical analyses included chi-square, McNemars, and logistic regression to assess accuracy, question fatigue, confidence calibration, and citation reliability. ResultsLLMs achieved high overall accuracy (78-87%), with the Individual Question format consistently yielding higher scores than Full Test, though differences were not statistically significant. Accuracy was highest in fact-dense domains (biochemistry, physiology, microbiology) and lowest in integrative domains (diagnosis, therapy). Significant question fatigue was observed in GPT-5.0 Full Test mode (OR = 0.997, p = 0.035), but not in Individual Question mode. Confidence scores predicted accuracy, with the strongest calibration in Individual Question mode. Citation analysis revealed frequent hallucinations, mostly critically erroneous, and citation validity was independent of answer accuracy. ConclusionsLLMs can answer a broad spectrum of periodontal specialty questions, but their reliability varies with context and information presentation. While promising as adjunctive tools, their outputs-- especially for complex reasoning and citations--require rigorous human review in educational and research settings to ensure accuracy and safety. Author SummaryArtificial intelligence chatbots are rapidly entering medical education, yet we lack comprehensive understanding of their reliability when students depend on them for learning. We developed a multidimensional evaluation framework to systematically assess AI performance beyond simple accuracy, examining how these systems behave across different medical topics, question types, and presentation formats. Using 331 real dental examination questions, we tested three major AI systems, analyzing not only correctness but also confidence calibration - whether AI confidence levels match actual accuracy - and implementing human-in-the-loop verification to check if cited sources actually exist. Our findings highlight critical vulnerabilities in current AI systems. Most alarmingly, these chatbots fabricated nearly half of their citations while maintaining unwavering confidence in both correct and incorrect responses. This combination of overconfidence and misinformation means students cannot distinguish reliable from unreliable AI responses. Additionally, we documented progressive performance decline during sequential questioning, similar to human cognitive fatigue. While we know AI systems generate rather than retrieve information, our research demonstrates the real-world consequences of this limitation. As artificial intelligence integrates into education, healthcare diagnostics, and insurance decisions, these findings underscore the urgent need for better evaluation frameworks and user education about AI limitations.
Rousseau, M.; Zouaq, A.; Huynh, N.
Show abstract
BackgroundThe near-exponential increase in the number of publications in orthodontics poses a challenge for efficient literature appraisal and evidence-based practice. Language models (LM) have the potential, through their question-answering fine-tuning, to assist clinicians and researchers in critical appraisal of scientific information and thus to improve decision-making. MethodsThis paper introduces OrthodonticQA (OQA), the first question-answering dataset in the field of dentistry which is made publicly available under a permissive license. A framework is proposed which includes utilization of PICO information and templates for question formulation, demonstrating their broader applicability across various specialties within dentistry and healthcare. A selection of transformer LMs were trained on OQA to set performance baselines. ResultsThe best model achieved a mean F1 score of 77.61 (SD 0.26) and a score of 100/114 (87.72%) on human evaluation. Furthermore, when exploring performance according to grouped subtopics within the field of orthodontics, it was found that for all LMs the performance can vary considerably across topics. ConclusionOur findings highlight the importance of subtopic evaluation and superior performance of paired domain specific model and tokenizer.
Mashhood, A.; Ahmed, A.; Khan, I.; Hashim, M.; Baloch, S.
Show abstract
BackgroundGAI tools are increasingly used informally for health, yet evidence from low- and middle-income countries (LMICs) is limited. This study generates early evidence on such health systems from the fifth most populous country: Pakistan. MethodsWe used a youth-led convergent mixed-methods design among digitally connected urban youth in Pakistan (survey N=1240, 20 interviews). The primary outcome was any GAI use for health. We fitted multivariable logistic regression models and conducted reflexive thematic analysis. FindingsOverall, 69.0% of participants reported using GAI for health. Higher odds of use were observed among women (aOR = 1.57, 95% CI [1.17-2.11], p = 0.003) and youth reporting any mental or physical condition (aOR = 1.82, 95% CI [1.34-2.48], p < .001). Greater trust in AI strongly predicted use (per-level aOR = 4.21, 95% CI [2.98-6.01], p < .001). High confidence using AI (aOR = 1.81, 95% CI [1.11-3.07], p = 0.022), awareness of AI risks (aOR = 1.67, 95% CI [1.20-2.31], p = 0.002), and prior use of other (non-generative) digital health tools (aOR = 4.48, 95% CI [2.59-8.23], p < .001) were also associated with higher likelihood of use. Telemedicine use was significant though weaker in magnitude (aOR = 1.58, 95% CI [1.01-2.54], p = 0.049). Interviews highlighted three themes: (1) access and affordability driving first-line use; (2) emotional safety and informational support, especially for stigmatized concerns; and (3) perceived empowerment in interpreting tests, organizing symptoms, and preparing for clinical visits. ConclusionGiven constrained, stigmatizing, and costly services, GAI may function as an adjunct "first step" for youth health information and emotional support in Pakistans health ecosystem.
Vumbugwa, P.; Puttkammer, N.; Majaha, M.; Stampfly, S.; Biondich, P.; Shivers, J. E.; Mburu, K.; Soge, O. O.; Longenecker, C.; Flowers, J.; Feldacker, C.
Show abstract
IntroductionCentral to a functional public health system is a strong health information ecosystem and robust data use. Many low-and-middle-income countries (LMICs) face the task of digitizing their health information systems (HIS). For health leaders, deciding what to prioritize when investing in HIS strengthening is central to this daunting challenge. ObjectivesThe study explores how HIS maturity assessment contributes to HIS strengthening, describes the facilitators and barriers to HIS maturity assessments, and how health leaders can prioritize conducting maturity assessments. MethodsThis descriptive qualitative study employed key informant interviews (KIIs) with fourteen eHealth leaders at national and international levels working or supporting Ministries of Healths national HIS in LMICs. Results were analyzed using Dedoose Version 9.0 to develop themes based on the health systems building blocks as a framework for identifying facilitators and barriers to conducting HIS maturity assessment. ResultsParticipants identified maturity assessments as a critical beginning step to HIS strengthening, showing the systems performance, and building a baseline response to systematic data quality challenges. Barriers to conducting HIS maturity assessment include lacking collaborators buy-in, fragmented vision, low financial/human resources, and overdependence on donor priorities. Non- supportive policies, a lack of execution champions, and an inadequately skilled workforce in conducting maturity assessments or negotiating for their prioritization hinder maturity assessment implementation. Frequently identified facilitators to promoting HIS maturity assessment include multi-stakeholder engagement, understanding the countrys HIS ecosystem, and priorities to appropriately integrate maturity assessment objectives. Recommendations include capacity building in data use and conducting maturity assessments at all health system levels to grow the demand and value of HIS maturity assessments. ConclusionPromoting HIS maturity assessments can help leaders prioritize areas to improve in the HIS ecosystem, making appropriate decisions that steward HIS maturity advancement. Addressing challenges that hinder HIS assessment implementation holds promise to identify a pathway to a strengthened health system. Author SummaryOur manuscript specifically spotlights the perspectives of African eHealth leaders, centering voices on the barriers and facilitators to planning and implementing HIS maturity assessments. We demonstrate their perspective on how conducting maturity assessments can inform understanding of gaps to address in the HIS and strategic direction. We detail the leaders recommendations for using HIS maturity assessments in strengthening HIS governance and overall health systems for better population health outcomes in LMIC settings.
Jain, D.; Rai, S.; Mittal, J.; Andy, A.; Buttenheim, A. M.; Guntuku, S. C.
Show abstract
BackgroundCOVID-19 vaccine hesitancy, fueled by concerns about vaccine development, side effects, and misinformation on social media platforms like Twitter, resulted in lower vaccination rates in Sub-Saharan Africa. MethodsWe collected, preprocessed, and geolocated 6,546,893 tweets related to COVID-19 vaccination from Sub-Saharan Africa. Using a vaccine misinformation classifier trained on RoBERTa embeddings, we identified 371,965 tweets in our dataset that included misinformation. We characterized the relationship between specific COVID-19 vaccine topics and the prevalence of misinformation, examined temporal variation in misinformation, and separately described the prevalence of misinformation in clusters defined by country-level socioeconomic and development metrics and by COVID-19 epidemiology. ResultsMisinformation in Sub-Saharan Africa is associated with discussions about pharmaceutical company profits, global access to vaccines and disparity, and trust in scientific research regarding vaccines. The prevalence of misinformation topics varied widely across country clusters as defined by socioeconomic development and COVID-19 epidemiology metrics. ConclusionsSocial media data provides valuable insights about vaccine hesitancy and vaccine misinformation in Sub-Saharan Africa that can inform policy and programmatic interventions to support vaccine demand and vaccine promotion.
Evans, L.; Evans, J.; Abdulla, A.; Ahmed, Z.
Show abstract
BackgroundDigital health has progressed rapidly due to the advances in technology and the promises of improved health and personal health empowerment. Concurrently, the burden of respiratory disease is increasing, particularly in Asia, where mortality rates are higher, and public awareness and government engagement are lower than in other regions of the world. Leveraging digital health interventions to manage and mitigate respiratory disease presents itself as a potentially effective approach. This study aims to undertake a scoping review to map respiratory digital health interventions in South and Southeast Asia, identify existing technologies, opportunities, and gaps, and put forward pertinent recommendations from the insights gained. MethodsThis study used a scoping review methodology as outlined by Arksey and OMalley and the Joanna Briggs Institute. Medline, Embase, CINAHL, PsycINFO, Cochrane Library, Web of Science, PakMediNet and MyMedR databases were searched along with key websites grey literature databases. ResultsThis scoping review has extracted and analysed data from 87 studies conducted in 14 South and Southeast Asian countries. Results were mapped to the WHO classification of digital health interventions categories to better understand their use. Digital health interventions are primarily being used for communication with patientes and between patients and providers. Moreover, interventions targeting tuberculosis were the most numerous. Many old interventions, such as SMS, are still being used but updated. Artificial intelligence and machine learning are also widely used in the region at a small scale. There was a high prevalence of pilot interventions compared to mature ones. ConclusionsThis scoping review collates and synthesises information and knowledge in the current state of digital health interventions, showing that there is a need to evaluate whether a pilot project is needed before starting, there is a need to report on interventions systematically to aid evaluation and lessons learnt, and that artificial intelligence and machine learning interventions are promising but should adhere to best ethical and equity practices. Author summaryTechnology has advanced quickly, facilitating the development of digital health, that is the use of technological tools for health purposes. Digital health tools may help more people achieve better health. At the same time, respiratory diseases are becoming a growing problem, especially in Asia, where there are more deaths and diseases linked to respiratory causes than in other parts of the world. Using digital health tools may be an effective way to manage and reduce the impact of respiratory diseases in the region. To that end, this study reviewed current digital health tools in South and Southeast Asia, identified gaps and opportunities and made recommendations based on the findings. The methodology used was a scoping review, which followed standards as described by Arksey and OMalley and the Joanna Briggs Institute. It searched relevant medical databases for information. This review includes 87 studies from 14 different countries. It revealed that tuberculosis was the most targeted disease by digital health interventions and that older technologies, such as the SMS, are still being used and updated as needed. Moreover, it revealed that new technologies like artificial intelligence and machine learning are being used more frequently but in small projects and that many of the projects described are small-scale pilot projects.
Wagon, L.; Nissen, M.; Kowatsch, T.
Show abstract
ObjectiveAs healthcare becomes increasingly digitalized, new challenges and opportunities arise for advancing health equity. In response, a broad range of frameworks and recommendations have emerged to guide equitable innovation. Yet despite this momentum, real world progress remains limited and many digital health organizations (DHOs) struggle to translate health equity principles into concrete, actionable strategies. One promising avenue for closing this implementation gap lies in the use of organizational readiness assessments that not only diagnose current capacity but also help to define a roadmap for improvement. In response, this study introduces the Health Equity Readiness Index and a supporting novel self-assessment tool designed to support DHOs in evaluating and advancing their readiness to act on health equity. MethodsDrawing on established frameworks from health equity, responsible AI, digital innovation, and organizational change, we created a five-dimensional Health Equity Readiness Index, encompassing strategy, governance, culture, data, and community collaboration. Each dimension was operationalized into 15 readiness indicators across four levels. A corresponding digital self-assessment tool was developed and refined through user testing and a structured online survey targeting DHOs. ResultsThe resulting tool provides a structured, low-barrier mechanism for DHOs to assess their current equity capacity and identify priority areas for improvement. Survey results (N = 124) showed broad applicability across diverse organizational contexts, with respondents spanning regions, service types, and roles. Participants rated the tool highly across all constructs (M [≥] 4.39/5). Qualitative feedback highlighted five overarching strengths (clarity, diagnostic value, actionability, useability, awareness) and three areas for future improvement (examples, customization, additional features). Overall, respondents perceived the tool as both relevant and actionable for guiding equity-oriented strategy. ConclusionThis study contributes a novel framework and diagnostic tool for advancing organizational readiness for health equity in digital health, supporting DHOs in moving from intention to implementation. It also offers potential utility for funders, regulators, and health systems seeking to institutionalize equity through incentives or benchmarking.
Tajudeen, R.; Fallah, M. P.; Ojo, J.; Shaweno, T.; Amanuel, W.; Sileshi, M.; Mulugeta, F.; Bamatura, M.; Kibiye, D.; Kabwe, P. C.; Sembuche, S. C.; Ngongo, N.; Dereje, N.; Kaseya, J.
Show abstract
The DHIS2 system enabled real-time tracking of vaccine distribution and administration to facilitate data-driven decisions. Experts from the Africa Centres for Disease Control and Prevention (Africa CDC) Monitoring and Evaluation (M&E) and Management Information System (MIS) teams, with support from the Health Information Systems Program South Africa (HISP-SA), developed the continental COVID-19 vaccination tracking system. Several variables related to COVID-19 vaccination were considered in developing the system. Three-hundred users can access the system at different levels with specific roles and privileges. Four dashboards with high-level summary visualizations were developed for top leadership for decision-making, while pages with detailed programmatic results are available to other users depending on their level of access. Africa CDC staff at different levels with a role-based account can view and interact with the dashboards and make necessary decisions based on the COVID-19 vaccination data from program implementation areas on the continent. The Africa CDC vaccination program dashboard provided essential information for public health officials to monitor the continental COVID-19 vaccination efforts and guide timely decisions. As the impact of COVID-19 is not yet over, the continental tracking of COVID-19 vaccine uptake and dashboard visualizations are used to provide the context of continental COVID-19 vaccination coverage and multiple other metrics that may impact the continental COVID-19 vaccine uptake. The lessons learned during the development and implementation of a continental COVID-19 vaccination tracking and visualization dashboard may be applied across various other public health events of continental and global concern. Author SummaryIn our work, we developed a real-time tracking system for COVID-19 vaccination across Africa, using the DHIS2 platform to help guide data-driven decisions. This system, created by experts from the Africa CDC in collaboration with the Health Information Systems Program South Africa (HISP-SA), enables us to monitor and manage COVID-19 vaccine distribution and administration continent-wide. With nearly 300 users from various levels of Africa CDC, the system allows authorized personnel to access relevant vaccination data tailored to their role, making the information easy to interpret and act upon. Our system includes four interactive dashboards that provide high-level summaries for top leadership and detailed views for program teams. These visual tools empower Africa CDC staff to make informed decisions, track progress, and address challenges in real-time. As COVID-19 remains a global concern, our platform provides crucial insights into vaccine coverage and related health metrics, demonstrating the potential for similar systems to enhance public health responses to future emergencies across Africa and beyond.
Pham, T.
Show abstract
Accurate classification of pediatric dental diseases from panoramic radiographs is crucial for early diagnosis and treatment planning. This study explores a text-based approach using a natural language transformer to generate textual descriptions of radiographs, which are then classified using deep learning models. Three models were evaluated: a one-dimensional convolutional neural network (1D-CNN), a long short-term memory (LSTM) network, and a pretrained bidirectional encoder representations from transformer (BERT) model for binary disease classification. Results showed that BERT achieved 77% accuracy, excelling in detecting periapical infections but struggling with caries identification. The 1D-CNN outperformed BERT with 84% accuracy, providing a more balanced classification, while the LSTM model achieved only 57% accuracy. Both 1D-CNN and BERT surpassed three pretrained CNN models trained directly on panoramic radiographs, indicating that text-based classification is a viable alternative to traditional image-based methods. These findings highlight the potential of language-based models for radiographic interpretation while underscoring challenges in generalizability. Future research should refine text generation, develop hybrid models integrating textual and image-based features, and validate performance on larger datasets to enhance clinical applicability.
Kuria, T.; Kamau, G.; Makokha, F.; Omondi, P.; Mbugua, G.; David, K.; Mbugua, S.; Gitaka, J.
Show abstract
Introduction: Timely, protocol-adherent clinical decisions are crucial for reducing neonatal mortality in low-resource settings. Translating extensive national guidelines into bedside practice remains challenging. Objective: We developed and evaluated AIFYA, a human-supervised, large language model LLM based clinical decision support system CDSS aligned with Kenya's national newborn care protocols. Methods: This prospective mixed methods early stage evaluation guided by the DECIDE-AI framework embedded AIFYA into routine workflows at two public health facilities Level 5 and Level 4 in Bungoma County Kenya from September 2024 to June 2025. Primary outcomes were adoption measured by cumulative neonatal cases managed training reach assessed by credentialed healthcare workers HCWs and guideline and citation concordance evaluated through blinded review of 118 AI generated recommendations by two neonatologists with adjudication by a third. Secondary outcomes included protocol adherence and triage to decision time. Results: A total of 50 HCWs were trained and 550 neonatal cases were managed over 10 months. Among surveyed HCWs n equals 33, 76 percent were female with mean age 32.1 years. Expert review found 75 percent of recommendations were correct and 15 percent partially correct with strong inter rater reliability weighted Cohen's kappa 0.85 and 95 percent CI 0.79 to 0.91. Citation accuracy was 96 percent. In 40 complex dosing scenarios 75 percent of outputs were rated correct. The median triage to decision time was 23 minutes with interquartile range 18 to 31. Implementation was supported by an offline first architecture and a facility based coaching model sustaining engagement despite staff turnover. Conclusion: A human supervised AI CDSS directly and transparently anchored to national clinical guidelines can be successfully implemented in routine low resource neonatal care settings. The system demonstrated high user adoption and strong expert rated concordance. High citation accuracy builds clinical trust ensuring safety and enabling auditable AI. These findings support progression to controlled multi site trials to evaluate clinical effectiveness. Keywords: Neonatal care Clinical decision support system Large language model Artificial intelligence Human supervised Low resource settings Guideline adherence Digital health Kenya
Uzochukwu, B. S. C.; Cherima, Y. J.; Enebeli, U. U.; Hassan, B.; Okeke, C. C.; Uzochukwu, A. C.; Omoha, A.; Uzochukwu, K. A.; Kalu, E. I.; Victor, D.; Alih, H. E.; Matinja, L. S.; Rindap, I. T.
Show abstract
Objective: To independently audit vendor-reported performance claims of health AI systems deployed in Nigeria and assess discrepancies, clinical consequences, equity impacts, and implications for safe AI deployment in low- and middle-income countries. Methods and analysis: We conducted a mixed-methods longitudinal audit (October 2024-March 2026) of six health AI systems (chest X-ray interpretation, TB screening, symptom triage, maternal health risk prediction, patient history intake, and health chatbots) across 73 diverse health facilities in six Nigerian states, involving 52,000 patients and 45 key informant interviews conducted with stakeholders. All data were sourced from integrated facility-level records, and no database linkage was performed. Vendor claims were abstracted from documentation, white papers, and validation studies. Independent performance was verified by an independent third party through system logs, patient records, clinical outcomes, and stakeholder interviews. Performance gaps were quantified as absolute percentage-point differences; clinical harms were estimated using patient volume and bootstrap confidence intervals; equity impacts were assessed across vulnerability dimensions (geography, age, income, comorbidities, infrastructure) using interaction terms in mixed-effects models and an Equity Harm Index (EHI). Results: Vendor-reported accuracy averaged 91.5%, while independently measured real-world accuracy averaged 67.3%, yielding a mean performance gap of 24.2 percentage points (95% CI: 21.5 to 26.9; p<0.001) across systems. Gaps ranged from 17 to 35 percentage points and were statistically significant for all systems. These discrepancies translated to substantial preventable harm, including an estimated 1,247 undetected TB cases (186 preventable deaths) and 342 misclassified high-risk pregnancies annually. Performance gaps were 28-38% larger among vulnerable groups (e.g., rural patients showed 38% higher EHI). Gaps were classified as systematic, context-dependent, or population-dependent. Conclusion: Vendor-reported performance metrics substantially overstated the real-world effectiveness of health AI in Nigeria, leading to preventable patient harm and widening inequities. Mandatory independent post-deployment verification, analogous to pharmaceutical Phase IV surveillance, is essential to ensure safe, equitable AI use in resource-constrained settings. Donors and regulators should prioritize verification over trust-based deployment.
D'addario, A. M. V.
Show abstract
The clinical promise of Large Language Models (LLMs) is often unrealized due to pro-hibitive computational costs. These costs create barriers not only to deployment in patient care but also to the vital process of fine-tuning models for specialized medical tasks and local patient populations. This study investigates 4-bit quantization as a methodology to make the entire clinical AI lifecycle--from development to implementation--both financially and practically viable. We performed a cost-benefit analysis using the Gemma 3 model family on the HealthQA-BR medical benchmark. We compared the diagnostic accuracy and computational resource requirements of standard full-precision models against their 4-bit quantized counterparts during both inference (clinical use) and QLoRA-based fine-tuning (model development). Quantization enabled massive efficiency gains with a clinically negligible impact on performance. For the 12B-parameter model, we observed a mere 1.3% absolute drop in accuracy. In exchange, computational requirements were reduced by 80% for fine-tuning and 69% for inference. This translates to a more than three-fold improvement in performance per unit of computational cost, accelerating research and development cycles. 4-bit quantization is a pivotal enabling technology for clinical AI. By drastically lowering the resource barrier for model customization and deployment, it empowers medical institutions to rapidly develop and validate specialized AI tools on-site. This approach holds particular promise for large-scale public health systems like Brazils SUS and provides a viable blueprint for similar health systems worldwide to transform AI from a theoretical possibility into a practical and equitable reality in patient care.
Dhodho, E.; Choga, K.; Mundoga, F.; Chimberengwa, P. T.; Gongora, R. T.; Webb, K.; Chinyanga, T. T.; Banda, F.; Masiye, K.; Midzi, N.; Mudavanhu, J.; Katsidzira, A.; Manyiyo, B.; Apollo, T.; Chimbetete, C.; Mhlanga, T.; Mangisi, P.; Gwanzura, C.; Tsvangirayi, S.; Dixon, J.; Nitsch, D.
Show abstract
Electronic health records (EHR) are increasingly recognised as critical digital infrastructure for integrated, patient-centred care in the context of rising multimorbidity. In low-resource settings, national EHRs may also support locally driven learning to improve adaptive care across chronic conditions. However, there is limited empirical evidence on whether and how these systems enable learning within routine care in ways that inform broader system adaptation. We conducted a qualitative multi-method assessment of Impilo, Zimbabwe's national EHR, to examine its capacity to support learning for integrated multimorbidity care at primary care level, using HIV-hypertension as a tracer condition pair. Guided by Friedman's socio-technical infrastructure model as the analytical framework and Learning Health Systems (LHS) theory as the interpretive framework, data were drawn from documentary review, ethnographic observation, patient journey mapping, and interviews with frontline health workers and key stakeholders. Frontline learning for person-centred multimorbidity care was actively generated through interpretation of patient trajectories, experiential adjustment, and coordination across HIV and hypertension services using both the EHR and paper-based artefacts such as registers and patient booklets. However, this learning remained largely encounter-bound and weakly stabilised. Impilo did not routinely provide usable longitudinal patient views, practice-facing analytic tools, or institutionalised mechanisms for collective reflection required to support integrated multimorbidity care. Consequently, learning was largely confined to incremental adjustment within existing workflows, with limited capacity to inform broader changes to care pathways, routines, or system design. These findings suggest that the principal barrier to developing LHS is not the absence of data or frontline learning capacity, but the lack of socio-technical arrangements that enable learning to stabilise and inform system adaptation. Digitalisation alone is insufficient to support adaptive multimorbidity care. Co-production with frontline health workers may provide a pathway for aligning digital system design with routine care realities.
Olatunji, T.; Aka, C.; Okocha, C.; Ayodele, E.; Orisakwe, J.; Adekunle, T.; Sanni, M.; Abiola, A.; Abdullahi, T.; Owopetu, O.; Afolaranmi, T.; Yougha, P. S.; Emmanuel-Fabula, M.; Menon, V.; Denniston, A.; Liu, X.; Williams, G.; Mateen, B. A.
Show abstract
In this study, we introduce a novel benchmark comprising over 9,000 real-world, point-of-care, multilingual, and multimodal clinical question-answer pairs sourced from frontline health workers in Nigeria. Using the dataset, we compare local general practitioners to multiple leading open and closed LLMs. Our results reveal several critical insights into the suitability of LLMs as clinical decision support systems in low-resource contexts. The results confirm that performance varies widely by language and input modality (e.g., text vs speech): while models perform best on English text inputs, their accuracy drops significantly for local-language speech. Critically, it is possible to achieve substantial performance gains by transcribing and translating other languages into English before prompting an LLM-- an important insight for non-anglophone product developers. Finally, this benchmark highlights key limitations of SLMs in supporting frontline healthcare in low-resource settings and provides a clear opportunity to track improvements as novel solutions are developed.
Moore, C.; Mugwagwa, J.; Vickers, I.
Show abstract
The use of Artificial Intelligence (AI) in healthcare is a field of growing relevance and importance, but in many LMICs, those seeking to develop AI based solutions for healthcare needs, face significant outstanding challenges. This research analysed practical efforts to implement AI-based technologies to support healthcare delivery in low-resource settings. By investigating six pilots within the Foreign Commonwealth and Development Offices Frontier Technologies program through analysis of associated pilot literature and semi-structured interviews with key pilot actors, we identified differences and commonalities in the experiences of each pilot, and in the perceived enablers and barriers for effective implementation of AI health tools. We found that AI is a promising tool in this sector but currently lacks the operating environment to be widely successful in solving healthcare challenges. Gaps in regulatory and ethical governance in these contexts exacerbated concerns around the ethical and responsible use of AI and led to alternative technical approaches being followed. The value of partnerships and relationships was demonstrated as essential, and projects with pre-established networks with key decision makers in healthcare systems, both at a bureaucratic and clinical level, demonstrated greater success in both developing and scaling their solutions. The challenge of sustainability and longer-term impact was also identified. The fragmented nature of local technology ecosystems also posed a common barrier to the delivery and scale-up of promising AI tools. It is anticipated that this research can help share some useful lessons for future users and developers of AI technologies and tools in the health space, particularly in resource-constrained settings. These findings suggest that barriers to equitable AI adoption in low-resource settings are primarily institutional and systemic, rather than technical, highlighting the need for health system-level readiness alongside technological innovation. Author SummaryArtificial intelligence (AI) is increasingly promoted as a way to improve healthcare delivery, including in low- and middle-income countries (LMICs). However, much of the existing discussion focuses on technical performance, with less attention to whether AI tools can be implemented, governed, and sustained within real-world health systems. In this study, we examine a set of AI-for-health pilot projects implemented in low-resource settings to understand what enables or constrains their adoption. Using interviews with practitioners and a review of project documentation, we explore how these pilots interacted with existing health system conditions, including workforce capacity, data infrastructure, governance arrangements, and institutional partnerships. We find that many of the challenges faced by AI projects are not primarily technical, but instead reflect broader system-level constraints, such as limited regulatory capacity, fragmented data systems, and reliance on external actors for development and maintenance. Our findings suggest that achieving equitable and inclusive AI for health requires more than developing effective technologies. It also requires sustained investment in the institutions, governance structures, and system capacities that allow AI tools to be safely adopted and integrated into health services. This study offers practical insights for policymakers, funders, and practitioners seeking to use AI in ways that strengthen health systems rather than bypass them.
Pham, T.
Show abstract
This study proposes a deep learning vision-language model for the automated diagnosis of pediatric dental diseases, with a focus on differentiating between caries and periapical infections. The model integrates visual features extracted from panoramic radiographs using methods of non-linear dynamics and textural encoding with textual descriptions generated by a large language model. These multimodal features are concatenated and used to train a 1D-CNN classifier. Experimental results demonstrate that the proposed model outperforms conventional convolutional neural networks and standalone language-based approaches, achieving high accuracy (90%), sensitivity (92%), precision (92%), and an AUC of 0.96. This work highlights the value of combining structured visual and textual representations in improving diagnostic accuracy and interpretability in dental radiology. The approach offers a promising direction for the development of context-aware, AI-assisted diagnostic tools in pediatric dental care.
Urli Hodges, E.; Crissman, K.; Chang, Z.; Ekeigwe, K.; Vissoci, J. R. N.; Udayakumar, K.
Show abstract
Community health workers are crucial to public health efforts in low- and middle-income countries. Despite their contributions to improving health, they encounter numerous barriers in performing their day-to-day work. Artificial intelligence applications may offer a potential solution to some of these barriers. We examined the implementation of smartphone-based artificial intelligence interventions with community health workers in Uganda, Rwanda, and Nigeria, including how the interventions were developed, how they were used, and successes and challenges of their implementation. This research identified four key considerations for intervention developers, implementing partners, and governments, with respect to implementation of these applications, including: 1) empowering community health workers through professionalization, compensation, and training; 2) understanding the digital ecosystem and aligning with digitization efforts; 3) designing the solution to fit local context, ensuring the development and training of the intervention is applicable to the local environment; and, 4) managing data responsibly with adherence to data privacy and security regulations. This research describes the opportunities for filling gaps in supervisory support, improving diagnosis, and making work more efficient. Providing the community health workforce with digital tools replaces onerous, inefficient paper-based record keeping, enables opportunities for improved supervision, and facilitates decision-making in settings where these workers may be the only point of access to health services and information. Author SummaryCommunity health workers are often the backbone of the health workforce in low- and middle-income countries. Despite their contributions to community health improvement, they often face barriers and inefficiencies in delivering services. The expansion of smartphone-based artificial intelligence interventions may offer solutions. We examined the implementation of smartphone-based artificial intelligence interventions with community health workers in Uganda, Rwanda, and Nigeria. We found that smartphone-based artificial intelligence applications may help to replace onerous, inefficient paper-based record keeping, enable opportunities for improved supervision, and facilitate decision-making in settings where these individuals may be the only health service providers. This research also identified four key considerations for intervention developers, implementing partners, and governments, with respect to implementation of these applications, including: 1) empowering community health workers through professionalization, compensation, and training; 2) understanding the digital ecosystem and aligning with digitization efforts; 3) designing the solution to fit local context, ensuring the development and training of the intervention is applicable to the local environment; and, 4) managing data responsibly with adherence to data privacy and security regulations.